Automatic post-synchronization of speech utterances
نویسنده
چکیده
The paper considers a prototype for automatic postsynchronization that consists of two basic components. As a first step, dynamic time warping is applied to compute the time-correspondence between an original utterance and an utterance that serves as the timing reference signal. In a second step, a time-scaling algorithm modifies the time structure of the original utterance accordingly. Informal diagnostic evaluation has shown that good results are obtained if the similarity between the acoustic-phonetic contents of the utterances is high. Possible ways for improving robustness against acoustic-phonetic differences, such as those that result from different coarticulation, are suggested.
منابع مشابه
Model-based Animation of Coverbal Gesture
Virtual conversational agents are supposed to combine speech with nonverbal modalities for intelligible and believeable utterances. However, the automatic synthesis of coverbal gestures still struggles with several problems like naturalness in procedurally generated animations, flexibility in pre-defined movements, and synchronization with speech. In this paper, we focus on generating complex m...
متن کاملAutomatic Feature Selection for Predicting Content of User Utterances in Dialogs
In task-oriented spoken dialog systems (SDS), the system often requests explicit confirmation of userprovided task-relevant concepts. The user utterance following a confirmation question (the postconfirmation utterance) is important to successful dialog outcomes. It may be a simple confirmation or rejection (e.g. yes, no, right, correct), or a correction or a topic change containing new concept...
متن کاملA Comparison between Three Methods of Language Sampling: Freeplay, Narrative Speech and Conversation
Objectives: The spontaneous language sample analysis is an important part of the language assessment protocol. Language samples give us useful information about how children use language in the natural situations of daily life. The purpose of this study was to compare Conversation, Freeplay, and narrative speech in aspects of Mean Length of Utterance (MLU), Type-token ratio (TTR), and the numbe...
متن کاملAutomatic Call Routing With Multiple Language Models
Our motivation is to perform call routing of utterances without recourse to transcriptions of the training data, which are very expensive to obtain. We therefore use phonetic recognition of utterances and search for salient phonetic sequences within the decodings. An important issue in phonetic recognition is the language model. It has been demonstrated [1] that the use of an iterative language...
متن کاملValidation of an Expressive Speech Corpus by Mapping Automatic Classification to Subjective Evaluation
This paper presents the validation of the expressive content of an acted corpus produced to be used in speech synthesis. The use of acted speech can be rather lacking in authenticity and therefore its expressiveness validation is required. The goal is to obtain an automatic classifier able to prune the bad utterances –with wrong expressiveness–. Firstly, a subjective test has been conducted wit...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 1997